Goto

Collaborating Authors

 complex data


Soft-ECM: An extension of Evidential C-Means for complex data

arXiv.org Artificial Intelligence

Clustering based on belief functions has been gaining increasing attention in the machine learning community due to its ability to effectively represent uncertainty and/or imprecision. However, none of the existing algorithms can be applied to complex data, such as mixed data (numerical and categorical) or non-tabular data like time series. Indeed, these types of data are, in general, not represented in a Euclidean space and the aforementioned algorithms make use of the properties of such spaces, in particular for the construction of barycenters. In this paper, we reformulate the Evidential C-Means (ECM) problem for clustering complex data. We propose a new algorithm, Soft-ECM, which consistently positions the centroids of imprecise clusters requiring only a semi-metric. Our experiments show that Soft-ECM present results comparable to conventional fuzzy clustering approaches on numerical data, and we demonstrate its ability to handle mixed data and its benefits when combining fuzzy clustering with semi-metrics such as DTW for time series data.


An Automated Data Mining Framework Using Autoencoders for Feature Extraction and Dimensionality Reduction

arXiv.org Artificial Intelligence

This study proposes an automated data mining framework based on autoencoders and experimentally verifies its effectiveness in feature extraction and data dimensionality reduction. Through the encoding-decoding structure, the autoencoder can capture the data's potential characteristics and achieve noise reduction and anomaly detection, providing an efficient and stable solution for the data mining process. The experiment compared the performance of the autoencoder with traditional dimensionality reduction methods (such as PCA, FA, T-SNE, and UMAP). The results showed that the autoencoder performed best in terms of reconstruction error and root mean square error and could better retain data structure and enhance the generalization ability of the model. The autoencoder-based framework not only reduces manual intervention but also significantly improves the automation of data processing. In the future, with the advancement of deep learning and big data technology, the autoencoder method combined with a generative adversarial network (GAN) or graph neural network (GNN) is expected to be more widely used in the fields of complex data processing, real-time data analysis and intelligent decision-making.


Reviews: Semi-crowdsourced Clustering with Deep Generative Models

Neural Information Processing Systems

A complex DGM is proposed that jointly models observations with crowdsourced annotations of whether or not two observations belong to the same cluster. This allows crowdsourcing non-expert annotations to help with clustering complex data. Importantly, the model is developed for the semi-supervised case, i.e., annotations are only observed for a small proportion of observation pairs. The authors propose a hierarchical VAE structure to model the observations, with a discrete latent-variable z \sim p(z \pi), a continuous latent variable x \sim p(x z), and observed data o \sim p(o x). This is paired with a two-coin David-Skene model which is conditioned on the mixture variable z for annotations: L \sim p(L z_i, z_j, \alpha, \beta), where \alpha and \beta are annotator-specific latent variables that model the "expertise" of the m_th annotator (precision and recall parameters, respectively). To the best of my understanding, through the dependence of the two-coin model on the latent mixture association, though it is not explicitly stated in the paper, z represents cluster association in the model.


Understanding the differences between AI, ML, and deep learning

#artificialintelligence

Artificial intelligence (AI), machine learning (ML), and deep learning (DL) are terms that are often used interchangeably. However, they are not the same thing. While they are all related to each other, they have different meanings and applications. In this article, we will explore the differences between AI, ML, and DL. AI is a broad term that encompasses all aspects of creating intelligent machines that can perform tasks that typically require human intelligence, such as recognizing speech, making decisions, and understanding natural language.


DeepSee.ai Inducted into JPMorgan Chase's Hall of Innovation

#artificialintelligence

DeepSee.ai, the creator and leading provider of Knowledge Process Automation (KPA), announced that it has been inducted into the JPMorgan Chase Hall of Innovation. The bank's Hall of Innovation award recognizes select emerging tech companies for their innovation, business value, and disruptive nature. "We're extremely honored to accept this award from JPMorgan Chase" "DeepSee has helped us automate manual post-trade checks supporting complex derivatives trading into AI-powered business outcomes," said Tom Damico, Global Head of Equities Operations, JPMorgan Chase. "We're already seeing efficiencies in post-trade processing and reconciliations, with more efficient deal review timeframes and more importantly, reduced operational risk." "We're extremely honored to accept this award from JPMorgan Chase," said Steve Shillingford, CEO of DeepSee.


Edge-Cloud Cooperation for DNN Inference via Reinforcement Learning and Supervised Learning

arXiv.org Artificial Intelligence

Deep Neural Networks (DNNs) have been widely applied in Internet of Things (IoT) systems for various tasks such as image classification and object detection. However, heavyweight DNN models can hardly be deployed on edge devices due to limited computational resources. In this paper, an edge-cloud cooperation framework is proposed to improve inference accuracy while maintaining low inference latency. To this end, we deploy a lightweight model on the edge and a heavyweight model on the cloud. A reinforcement learning (RL)-based DNN compression approach is used to generate the lightweight model suitable for the edge from the heavyweight model. Moreover, a supervised learning (SL)-based offloading strategy is applied to determine whether the sample should be processed on the edge or on the cloud. Our method is implemented on real hardware and tested on multiple datasets. The experimental results show that (1) The sizes of the lightweight models obtained by RL-based DNN compression are up to 87.6% smaller than those obtained by the baseline method; (2) SL-based offloading strategy makes correct offloading decisions in most cases; (3) Our method reduces up to 78.8% inference latency and achieves higher accuracy compared with the cloud-only strategy.


Data Science Vs. Machine Learning: What's The Difference

#artificialintelligence

Data Science and Machine Learning are two different approaches to data analysis. Machine learning is a subset of data science and focuses more on training models and specific tasks. At the same time, Data Science encompasses more of the broader approach, analyzing large amounts of data across many fields. This article compares the two, highlighting their differences and areas they overlap. Data Science is the practice of extracting value from data to improve business outcomes.


The Power Of Machine Learning In Education Sector

#artificialintelligence

In the last few years, machine learning (ML) has been making some giant leaps in education – from predicting the next steps students need to take to improve their grades to generating teacher study material. This article discusses how machine learning can be used for education in more detail and some of the current trends in this field. A machine learning branch of artificial intelligence employs algorithms to learn from data. It can improve the accuracy, speed, and efficiency of various tasks, such as predicting customer behavior or organizing data. In the education sector, machine learning can help teachers identify and diagnose problems with their student's academic progress and help them decide which courses to teach. Machine learning can also be used to develop educational programs that can adapt to the needs of individual students.


SparkBeyond Discovery Now Available in the Microsoft Azure Marketplace

#artificialintelligence

SparkBeyond announced the availability of its data science platform for supervised machine learning, SparkBeyond Discovery, in the Microsoft Azure Marketplace, an online store providing applications and services for use on Azure. SparkBeyond customers can now take advantage of the productive and trusted Azure cloud platform, with streamlined deployment and management. SparkBeyond Discovery is a data science platform for supervised machine learning that helps data professionals save time, deepen their understanding of the problem space and improve model performance by automating feature discovery in complex data. The platform discovers known and unknown features in data that give information about a target to save data professionals time in feature engineering, and to maximize data value through applying a wide range of aggregations. The platform finds expressive, interpretable and concise features ranking them based on their predictive power and then applies AutoML to build glass-box models.


More complex deep learning models require more complex data

#artificialintelligence

This Massachusetts Institute of Technology (MIT) recent paper shows a more than interesting approach towards how machines can get to understand and interpret the relationships between objects in a scene. As this and other recent studies reflect, the pain is latent… deep learning models are getting very good at identifying objects in all kinds of scenes, however they can't understand the relationships of those objects with each other and the surrounding environment. Even simple relationships that are obvious for a human, like this is inside of this or that is on top of that, are very hard for widely used object detection and segmentation models. There is a growing number of use cases that will require this understanding. This evolution in the models will require new data for training.